Papers with audio-language models
SoundMind: RL-Incentivized Logic Reasoning for Audio-Language Models (2025.emnlp-main)
Copied to clipboard
Xingjian Diao, Chunhui Zhang, Keyi Kong, Weiyi Wu, Chiyu Ma, Zhongyu Ouyang, Peijun Qing, Soroush Vosoughi, Jiang Gui
| Challenge: | Recent large language models have demonstrated impressive reasoning abilities, but their extension to the audio modality remains underexplored. |
| Approach: | They propose a rule-based reinforcement learning algorithm to equip LALMs with robust reasoning capabilities. |
| Outcome: | The proposed algorithm improves on the SoundMind benchmark. |
AIR-Bench: Benchmarking Large Audio-Language Models via Generative Comprehension (2024.acl-long)
Copied to clipboard
Qian Yang, Jin Xu, Wenrui Liu, Yunfei Chu, Ziyue Jiang, Xiaohuan Zhou, Yichong Leng, Yuanjun Lv, Zhou Zhao, Chang Zhou, Jingren Zhou
| Challenge: | Existing benchmarks for audio-centric interaction have impeded advancements in this field . AIR-Bench evaluates LALMs' ability to understand audio signals and interact with humans . |
| Approach: | They propose a benchmark to evaluate the ability of large audio-language models to understand audio signals . they use 19 tasks with approximately 19k single-choice questions to examine single-task ability . |
| Outcome: | The proposed framework evaluates the ability of large audio-language models to understand audio signals and interact with humans in the textual format. |
Jamendo-MT-QA: A Benchmark for Multi-Track Comparative Music Question Answering (2026.findings-acl)
Copied to clipboard
Junyoung Koh, Jaeyun Lee, Soo Yong Kim, Gyu Hyeong Choi, Jung In Koh, Jordan Phillips, Yeonjin Lee, Min Song
| Challenge: | Existing benchmarks for music question answering do not systematically evaluate reasoning across tracks. |
| Approach: | They propose a dataset and benchmark for multi-track comparative question answering . they construct 36,519 comparative QA items over 12,173 track pairs . |
| Outcome: | The proposed dataset and benchmark for multi-track comparative question answering is based on the Jamendo-QA dataset. |
Unlocking Large Audio-Language Models for Interactive Language Learning (2026.findings-eacl)
Copied to clipboard
| Challenge: | Computer-Assisted Pronunciation Training (CAPT) systems provide unintuitive feedback that lacks actionable guidance. |
| Approach: | They propose to use audio-language models to provide more user-friendly feedback for pronunciation training. |
| Outcome: | The proposed model outperforms baselines on mispronunciation detection and suggestion generation. |
FineLAP: Taming Heterogeneous Supervision for Fine-grained Language-Audio Pretraining (2026.acl-long)
Copied to clipboard
| Challenge: | Existing audio-language models excel at clip-level understanding but struggle with frame-level tasks. |
| Approach: | They propose a novel training paradigm that advances both clip- and frame-level alignment in CLAP with heterogeneous data. |
| Outcome: | The proposed training paradigm improves both clip- and frame-level alignment in CLAP with heterogeneous data. |
Discovering and Causally Validating Emotion-Sensitive Neurons in Large Audio-Language Models (2026.acl-long)
Copied to clipboard
| Challenge: | Emotion is a central dimension of spoken communication, yet we lack a mechanistic account of how LALMs encode it internally. |
| Approach: | They propose to use emotion-sensitive neurons in large audio-language models to study their interpretations. |
| Outcome: | The proposed models show that they can be used to make decisions on emotion . the results show that the ESNs exhibit non-uniform clustering with partial cross-dataset transfer . |
EMO-RL: Emotion-Rule-Based Reinforcement Learning Enhanced Audio-Language Model for Generalized Speech Emotion Recognition (2025.findings-emnlp)
Copied to clipboard
| Challenge: | Recent advances in reinforcement learning (RL) have shown promise in improving LALMs’ reasoning abilities, but their performance in affective computing tasks remains suboptimal. |
| Approach: | They propose a framework incorporating reinforcement learning with two key innovations: Emotion Similarity-Weighted Reward (ESWR) and Explicit Structured Reasoning (ESR). |
| Outcome: | The proposed framework improves LALMs' reasoning abilities on MELD and IEMOCAP datasets and shows strong generalization. |
iKnow-audio: Integrating Knowledge Graphs with Audio-Language Models (2025.emnlp-main)
Copied to clipboard
| Challenge: | Contrastive language-audio pretraining models learn by aligning audio and text in a shared embedding space. |
| Approach: | They propose a framework that integrates knowledge graphs with audio-language models to provide robust semantic grounding. |
| Outcome: | iKnow-audio improves disambiguation of acoustically similar sounds and reduces prompt engineering. |